Papers with Handwritten Text Recognition
Books of Hours. the First Liturgical Data Set for Text Segmentation. (2020.lrec-1)
Copied to clipboard
Amir Hazem, Beatrice Daille, Christopher Kermorvant, Dominique Stutzmann, Marie-Laurence Bonhomme, Martin Maarand, Mélodie Boillet
| Challenge: | Until now, the book of hours has been scarcely studied because of its manuscript nature, its length and its complex content. |
| Approach: | They propose to use Handwritten Text Recognition to generate a corpus of Latin transcriptions of 300 books of hours generated by OCR for handwritten and not printed texts. |
| Outcome: | The proposed structure and state-of-the-art methods are compared with existing methods and are based on the results of a systematic evaluation of two books of hours. |
How Much Data Do You Need? About the Creation of a Ground Truth for Black Letter and the Effectiveness of Neural OCR (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in Optical Character Recognition and Handwritten Text Recognition have led to more accurate text recognition of historical documents. |
| Approach: | They propose to build a ground truth for a German-language newspaper published in black letter . they also evaluate the performance of different OCR engines and estimate how much data is needed to achieve high-quality OCR results. |
| Outcome: | The proposed model can recognise black letter text and performs well on data they have not seen during training. |
Evaluation of HTR models without Ground Truth Material (2022.lrec-1)
Copied to clipboard
Phillip Benjamin Ströbel, Martin Volk, Simon Clematide, Raphael Schwitter, Tobias Hodel, David Schoch
| Challenge: | Optical Character Recognition (OCR) is a well-established technique for digitising historical printed collections in libraries and archives. |
| Approach: | They propose to use masked language models to evaluate handwritten text recognition models . they propose to introduce GT-free metrics to evaluate models to ensure best results . |
| Outcome: | The proposed model evaluations are based on lexicon-based and masked language models. |
Digitizing Nepal’s Written Heritage: A Comprehensive HTR Pipeline for Old Nepali Manuscripts (2026.acl-long)
Copied to clipboard
| Challenge: | Using a line-level transcription approach, we explore encoder-decoder architectures and data-centric techniques to improve recognition accuracy for Old Nepali manuscripts. |
| Approach: | They propose a line-level transcription approach and explore encoder-decoder architectures and data-centric techniques to improve recognition accuracy. |
| Outcome: | The proposed model achieves a 4.9% error rate and is highly reliable. |
Automatic Transcription of Handwritten Old Occitan Language (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to handwritten text recognition have shown promising results, but low-resource languages often lack resources. |
| Approach: | They propose an HTR approach that leverages the Transformer architecture for recognizing handwritten Old Occitan language. |
| Outcome: | The proposed approach surpasses state-of-the-art models for Old Occitan HTR, including open-source Transformer-based models and commercial applications like Google Cloud Vision. |